Papers with Japanese language
Controlling Japanese Honorifics in English-to-Japanese Neural Machine Translation (D19-52)
Copied to clipboard
| Challenge: | In the Japanese language different levels of honorific speech are used to convey respect, deference, humility, formality and social distance. |
| Approach: | They propose a method for controlling the level of formality of Japanese output . they use heuristics to identify honorific verb forms to classify Japanese sentences . |
| Outcome: | The proposed model can produce Japanese translations in different honorific speech styles for the same English input sentence. |
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation (2025.naacl-long)
Copied to clipboard
Shota Onohara, Atsuyuki Miyai, Yuki Imajuku, Kazuki Egashira, Jeonghun Baek, Xiang Yue, Graham Neubig, Kiyoharu Aizawa
| Challenge: | Using culture-agnostic subsets, performance drops in many LMMs when evaluated in Japanese. |
| Approach: | They introduce a Japanese benchmark to evaluate large multimodal models on expert-level tasks based on the Japanese cultural context. |
| Outcome: | The proposed benchmark enables comparisons with other benchmarks in other languages based on cultural contexts. |
Topicalization in Language Models: A Case Study on Japanese (2022.coling-1)
Copied to clipboard
| Challenge: | a recent study has shown that neural language models can capture discourse-level preferences in text generation . a particular aspect of discourse is the topic-comment structure . |
| Approach: | They analyze whether neural language models can capture discourse-level preferences in text generation . they use Japanese language and crowdsourced human topicalization judgment data . |
| Outcome: | The proposed model can capture human-like generalizations in discourse-level linguistic aspects. |
Simplified Corpus with Core Vocabulary (L18-1)
Copied to clipboard
| Challenge: | a study has found that simple Japanese is more accessible to foreigners than English. |
| Approach: | They have constructed a simplified corpus for the Japanese language and selected the core vocabulary. |
| Outcome: | The simplified corpus can be used for automatic text simplification and translating simple Japanese into English and vice-versa. |
Universal Dependencies Version 2 for Japanese (L18-1)
Copied to clipboard
Masayuki Asahara, Hiroshi Kanayama, Takaaki Tanaka, Yusuke Miyao, Sumire Uematsu, Shinsuke Mori, Yuji Matsumoto, Mai Omura, Yugo Murawaki
| Challenge: | UD Japanese resources are built on automatic conversion from several treebanks. |
| Approach: | They propose to port the word delimitation, POS, and syntactic relations of existing treebanks to UD Japanese . they discuss the issues of the UD scheme found through porting of the Japanese language . |
| Outcome: | The proposed UD Japanese resources are based on automatic conversion from treebanks. |
Detecting Sensitive Personal Information in Japanese Pre-Training Corpora for Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Large-scale pre-training corpora are essential for large language models, but if such content remains unfiltered, there is a risk that LLMs may memorize it and leak it through their outputs. |
| Approach: | They construct a Japanese text corpora dataset and train machine learning models to detect SCPI in text. |
| Outcome: | The proposed classifier can detect information related to SCPI in Japanese text. |
A Large-Scale Japanese Dataset for Aspect-based Sentiment Analysis (2022.lrec-1)
Copied to clipboard
| Challenge: | Aspect-based sentiment analysis (ABSA) has not been explored in the Japanese language . there is no standard Japanese dataset available for ABSA task in the language - a paper by cnn. |
| Approach: | They propose to use a Japanese aspect-based sentiment analysis dataset for hotel reviews domain . they propose to include 53,192 review sentences with seven aspect categories and two polarity labels . |
| Outcome: | The proposed dataset contains 53,192 review sentences with seven aspect categories and two polarity labels. |
JLBert: Japanese Light BERT for Cross-Domain Short Text Classification (2024.lrec-main)
Copied to clipboard
| Challenge: | Short Texts face the problem of being short, equivocal, and non-standard. |
| Approach: | They propose a Japanese BERT model with cross-domain functionality and comparable accuracy to State of the Art models. |
| Outcome: | The proposed model outperforms state-of-the-art models on three short text datasets by 1.5% across various domains. |